‹ BackNewsrobot learning

robot learning

World Models
2026-09-11 02:56:12

What counts as a world model? Fei-Fei Li, LeCun and Zhu Jun offer three different answers

The term "world model" has become one of the least settled concepts in AI in 2026. It can refer to a model that generates coherent video, a digital environment that keeps responding to user input, a latent-space system that predicts future states, or a policy that outputs robot actions. Those systems all process information about the world, yet they do not solve the same problem, and comparing them by visual quality, geometric consistency, prediction accuracy or task success can blur more than it clarifies. A MarsBit article, citing work from World Labs, Meta and Tsinghua University professor Zhu Jun’s team, lays out three leading interpretations. Fei-Fei Li and World Labs sort world models into renderers, simulators and planners. Yann LeCun argues for learning predictable structure in an abstract latent space. Zhu Jun and his collaborators define a broader "general world model" from first principles, built around three linked capabilities: understanding, imagination and action. Their framework also proposes a five-level roadmap from L1 world generation to L5 world organization, and places video generation, real-time interaction and embodied control on a single capability curve. The piece further examines a data pyramid spanning internet video to robot interaction, the MoT architecture for multimodal coordination, and Motus2, which combines policy generation, simulation, evaluation and tactile feedback in one closed-loop system.

840
What counts as a world model? Fei-Fei Li, LeCun and Zhu Jun offer three different answers
Agibot removes chief scientist Luo Jianlan from partner roster as departure speculation grows